Terrorism
Democracy v the machine: the birth of the digital age and the warnings that were ignored
A women uses the new IBM 650 Magnetic Drum Data Processing Machine in a New York office in 1954. A women uses the new IBM 650 Magnetic Drum Data Processing Machine in a New York office in 1954. Many hoped that the march of technology would usher in an egalitarian utopia - but some foresaw the threat it would pose to liberal society. One of the stranger things about this dizzying, headlong moment in time is that it doesn't have much of a past. Everything is about the future of this, the future of that: the future of work, the future of humanity, the future of the planet. It's as if everyone is screaming (some ecstatically, most terror-stricken): robots are taking over the world! Meanwhile, you can't put down your phone, unplug, delete your AI apps, tell Zoom to piss off; it feels as if you are racing toward something, and can't stop, or look back, or think straight. But of course this weird moment in history does have a past. Things could have turned out differently.
The Key to Understanding the Disturbing Relationship Between Silicon Valley and the Trump Administration
Gil Durรกn's new book,, catalogs how tech oligarchs like Peter Thiel took over America. Enter your email to receive alerts for this author. You can manage your newsletter subscriptions at any time. You're already subscribed to the aa_Laura_Miller newsletter. You can manage your newsletter subscriptions at any time.
'If we don't fight back, we don't have a future': the journalist taking on the 'tech fascists' of Silicon Valley
'Maga is a cult of grievance that seeks to purge its enemies by demonising them through absurd and surreal propaganda' Gil Durรกn. 'Maga is a cult of grievance that seeks to purge its enemies by demonising them through absurd and surreal propaganda' Gil Durรกn. 'If we don't fight back, we don't have a future': the journalist taking on the'tech fascists' of Silicon Valley I n April, Gil Durรกn was permanently banned from Elon Musk's X for posting just two words. Durรกn was responding to a post from the tech company Palantir outlining its 22 point "technological manifesto", which praised American "hard power", western culture and AI weapons, denounced inclusivity and called for compulsory national service. Durรกn responded: "TLDR: Fascism" (TLDR is short for "too long, didn't read").
The Download: AI hiring biases, and weather data sabotage
Plus: SpaceX is negotiating to sell the Pentagon AI compute. The next time you apply for a job, AI may screen your rรฉsumรฉ before any human sees it. But there's good reason to question whether AI will judge you fairly. We already know that LLMs pick up human biases from their training data. New research suggests they can also develop their own biases from experience--and stereotype job applicants more than humans do. As AI companies race to build agentic models that remember the tiniest details about users, they may be handing them ammunition for forming those biases.
Trump Moves to Revoke Syria's Designation as State Sponsor of Terrorism
Follow this section to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. Follow this tag to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens.
4d18c7389f436e1e22b219d7e8d43f94-Paper-Conference.pdf
Alignment faking in large language models presented a demonstration of Claude 3 Opus and Claude 3.5 Sonnet selectively complying with a helpfulonly training objective to prevent modification of their behavior outside of training. We expand this analysis to 25 models and find that only 5 (Claude 3 Opus, Claude 3.5 Sonnet, Llama 3 405B, Grok 3, Gemini 2.0 Flash) comply with harmful queries more when they infer they are in training than when they infer they are in deployment. First, we study the motivations of these 5 models. Results from perturbing details of the scenario suggest that only Claude 3 Opus's compliance gap is primarily and consistently motivated by trying to keep its goals. Second, we investigate why many chat models don't fake alignment. Our results suggest this is not entirely due to a lack of capabilities: many base models fake alignment some of the time, and post-training eliminates alignment-faking for some models and amplifies it for others.We investigate 5 hypotheses for how post-training may suppress alignment faking and find that variations in refusal behavior may account for a significant portion of differences in alignment faking.
CTRL-ALT-DECEIT Sabotage Evaluations for Automated AIR&D
AI systems are increasingly able to autonomously conduct realistic software engineering tasks, and may soon be deployed to automate machine learning (ML) R&D itself. Frontier AI systems may be deployed in safety-critical settings, including to help ensure the safety of future systems. Unfortunately, frontier and future systems may not be sufficiently trustworthy, and there is evidence that these systems may even be misaligned with their developers or users. Therefore, we investigate the capabilities of AI agents to act against the interests of their users when conducting ML engineering, by sabotaging ML models, sandbagging their performance, and subverting oversight mechanisms. First, we extend MLE-Bench, a benchmark for realistic ML tasks, with code-sabotage tasks such as implanting backdoors and purposefully causing generalisation failures.
Reductio 1k ff((ff (+ฮปฮปฮปhhฮธฮธฮธฮธ1k1kXXYk12((((((H+i ii, estima Scientific study MUG'''3212'''223302222
Randomized experiments are the preferred approach for evaluating the effects of interventions, but they are costly and often yield estimates with substantial uncertainty. On the other hand, in silico experiments leveraging foundation models offer a cost-effective alternative that can potentially attain higher statistical precision. However, the benefits of in silico experiments come with a significant risk: statistical inferences are not valid if the models fail to accurately predict experimental responses to interventions. In this paper, we propose a novel approach that integrates the predictions from multiple foundation models with experimental data while preserving valid statistical inference. Our estimator is consistent and asymptotically normal, with asymptotic variance no larger than the standard estimator based on experimental data alone. Importantly, these statistical properties hold even when model predictions are arbitrarily biased. Empirical results across several randomized experiments show that our estimator offers substantial precision gains, equivalent to a reduction of up to 20% in the sample size needed to match the same precision as the standard estimator based on experimental data alone.
903ceb0ed2d5ceec6e2c9b317b6c54a8-Paper-Conference.pdf
Recent advances in Large Vision-Language Models (LVLMs) have showcased strong reasoning abilities across multiple modalities, achieving significant breakthroughs in various real-world applications. Despite this great success, the safety guardrail of LVLMs may not cover the unforeseen domains introduced by the visual modality. Existing studies primarily focus on eliciting LVLMs to generate harmful responses via carefully crafted image-based jailbreaks designed to bypass alignment defenses. In this study, we reveal that a safe image can be exploited to achieve the same jailbreak consequence when combined with additional safe images and prompts. This stems from two fundamental properties of LVLMs: universal reasoning capabilities and safety snowball effect. Building on these insights, we propose Safety Snowball Agent (SSA), a novel agent-based framework leveraging agents' autonomous and tool-using abilities to jailbreak LVLMs. SSAoperates through two principal stages: (1) initial response generation, where tools generate or retrieve jailbreak images based on potential harmful intents, and (2) harmful snowballing, where refined subsequent prompts induce progressively harmful outputs. Our experiments demonstrate that SSAcan use nearly any image to induce LVLMs to produce unsafe content, achieving high success jailbreaking rates against the latest LVLMs. Unlike prior works that exploit alignment flaws, SSAleverages the inherent properties of LVLMs, presenting a profound challenge for enforcing safety in generative multimodal systems.